image second
AI Chip Startup Puts Inference Cards on the Table
In the deep learning inferencing game, there are plenty of chipmakers, large and small, developing custom-built ASICs aimed at this application set. But one obscure company appears to have beat them to the punch. Habana Labs, a fabless semiconductor startup, began sampling its purpose-built inference processor for select customers back in September 2018, coinciding with the company's emergence from stealth mode. Eitan Medina, Habana's Chief Business Officer, claims its HL-1000 chip is now "the industry's highest performance inference processor." It's being offered to customers in a PCIe card that goes by name of Goya.
Lies, Damn Lies, And TOPS/Watt
There are almost a dozen vendors promoting inferencing IP, but none of them gives even a ResNet-50 benchmark. The only information they state typically is TOPS (Tera-Operations/Second) and TOPS/Watt. These two indicators of performance and power efficiency are almost useless by themselves. So what, exactly, does X TOPS really tell you about performance for your application? When a vendor says their ABC-inferencing-engine does X TOPS, you would assume that in one second it will perform X Trillion Operations.
Move Over GPUs: Startup's Chip Claims to Do Deep Learning Inference Better
Habana Labs, a startup that came out of "stealth mode" this week, announced a custom chip that is said to enable much higher machine learning inference performance compared to GPUs. According to the startup, its Goya chip is designed from scratch for deep learning inference, unlike GPUs or other types of chips that have been repurposed for this task. The chip's die is composed of eight VLIW Tensor Processing Cores (TPCs), each having their own local memory, as well as access to shared memory. The external memory is accessed through a DDR4 interface. The Goya chip supports all the major machine learning software frameworks, including TensorFlow, MXNet, Caffe2, Microsoft Cognitive Toolkit, PyTorch and the Open Neural Network Exchange Format (ONNX).
30,000 Images/Second: Xilinx and AMD Claim AI Inferencing Record
On the heels of its dual announcement at the Open Compute Project Summit in Amsterdam this week (see related story), Xilinx yesterday disclosed that AMD and Xilinx have teamed to set an AI inference processing record of 30,000 images per second. The joint work of the two companies, announced at the Xilinx Developer Forum in San Jose by Xilinx CEO Victor Peng and AMD CTO Mark Papermaster, connects AMD's EPYC CPUs and the new Xilinx Alveo FPGA accelerator card, announced yesterday at the OCP Summit. The record, running a batch size of 1 and Int8 precision, was accomplished on a system that leverages two AMD EPYC 7551 server CPUs with PCIe connectivity, along with eight Alveo U250 accelerator cards. In a blog post, Xilinx said the inference performance is powered by Xilinx ML Suite, which allows developers to optimize and deploy accelerated inference and supports various machine learning frameworks, such as TensorFlow. The benchmark was performed on the GoogLeNet convolutional neural network.